Papers with visual fidelity

13 papers
TeachMaster: Generative Teaching via Code (2026.acl-industry)

Copied to clipboard

Challenge: Existing methods for creating video content are limited by high costs and slow update cycles.
Approach: They propose a paradigm shifting educators from manual creators to high-level directors who focus on pedagogical intents while agents handle execution.
Outcome: The proposed framework reduces production costs to 0.3% of traditional course videos and provides a robust solution for scalable education.
Mirror in the Model: Ad Banner Image Generation via Reflective Multi-LLM and Multi-modal Agents (2025.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in generative modeling have greatly improved image synthesis quality.
Approach: They propose an agentic refinement framework for automatic ad banner generation that integrates a hierarchical multimodal agent system with a coordination loop.
Outcome: The proposed model outperforms existing models in real-world banner design scenarios.
FrontCoder: Scaling Visual Fidelity in Front-End Code Generation (2026.findings-acl)

Copied to clipboard

Challenge: Existing work on front-end code generation fails to provide visual fidelity and rendering quality for front- end developers.
Approach: They propose a three-stage pipeline to enhance front-end code generation capabilities in LLMs . they use synthetic data, quality-controlled supervised fine-tuning, and reinforcement learning .
Outcome: The proposed model achieves competitive performance with frontier models while maintaining generation efficiency.
Natural Language Rationales with Full-Stack Visual Reasoning: From Pixels to Semantic Frames to Commonsense Graphs (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use natural language rationales provide intuitive, higher-level explanations that are easily understandable by humans.
Approach: They propose a model that generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Outcome: The proposed model generates free-text rationales by combining pretrained language models with object recognition, grounded visual semantic frames, and visual commonsense graphs.
Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance (2026.eacl-long)

Copied to clipboard

Challenge: Existing methods for image-to-image translation lack structural integrity and attribute-specific control . Existing approaches lack semantics and provide fine-grained, attribute-based control compared to GAN-based methods .
Approach: They propose a language-grounded attribute-controllable translation framework that grounds semantic differences into corresponding visual transformations while preserving unrelated structural and semantic content.
Outcome: Experiments on CelebA(Dialog) and BDD100K show that LACE achieves high visual fidelity, structural preservation, and interpretable domain-specific control, surpassing baselines.
Evian: Towards Explainable Visual Instruction-tuning Data Auditing (2026.findings-acl)

Copied to clipboard

Challenge: Existing data filtering methods rely on coarse-grained scores that lack granularity to identify nuanced semantic flaws.
Approach: They propose a "Decomposition-then-Evaluation" paradigm that breaks model responses into constituent cognitive components.
Outcome: The proposed model outperforms models trained on larger datasets in three key areas . the authors show that Logical Coherence is the most critical factor in data quality evaluation .
Diffusion-CAM: Faithful Visual Explanations for dMLLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing Class Activation Mapping methods are ill-suited for interpreting non-autoregressive behaviors of diffusion-based architectures.
Approach: They propose to use a method to generate parallel activation maps by probing intermediate representations in the transformer backbone to capture latent features and their class-specific gradients.
Outcome: Experiments show that Diffusion-CAM significantly outperforms SoTA methods in localization accuracy and visual fidelity.
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage (2026.acl-long)

Copied to clipboard

Challenge: Vision-Language Models excel at photorealistic generation, but struggle to represent abstract meanings.
Approach: They propose a benchmark that replaces high-fidelity visual detail with schematic iconicity by generating paired, sense-anchored visualizations for literal and idiomatic readings.
Outcome: The proposed benchmark replaces high-fidelity visual detail with schematic iconicity by generating paired, sense-anchored visualizations for literal and idiomatic readings.
Efficient Inference for Large Vision-Language Models: Bottlenecks, Techniques, and Prospects (2026.findings-acl)

Copied to clipboard

Challenge: Large Vision-Language Models are hindered by a systemic efficiency barrier known as visual token dominance.
Approach: They propose a systematic taxonomy of efficiency techniques structured around the inference lifecycle . they examine visual encoding, prefilling, and decoding to understand bottlenecks .
Outcome: The proposed techniques reveal how upstream decisions dictate downstream bottlenecks . the proposed techniques include hybrid compression and modality-aware decoding .
Automatic Evaluation for Text-to-image Generation: Task-decomposed Framework, Distilled Training, and Meta-evaluation Benchmark (2025.acl-long)

Copied to clipboard

Challenge: Existing MLLMs rely on commercial models such as GPT-4o for evaluations, but they are not universally accessible.
Approach: They propose a task decomposition evaluation framework based on GPT-4o to automatically construct a specialized training dataset to break down the multifaceted evaluation process into simpler sub-tasks.
Outcome: The proposed framework outperforms the current state-of-the-art GPT-4o evaluation framework with over 4.6% improvement in Spearman and Kendall correlations with human judgments.
PlotGen-Bench: Evaluating VLMs on Generating Visualization Code from Diverse Plots across Multiple Libraries (2026.findings-acl)

Copied to clipboard

Challenge: PlotGen-Bench evaluates vision-language models' ability to generate executable visualization code from plots under realistic and complex visualization requirements.
Approach: They propose a benchmark to evaluate plot-to-code generation in vision-language models . they use Matplot, Matplos, Mat3D, Mat4D, and Mat4E to evaluate their performance .
Outcome: The proposed benchmark covers 9 major categories, 30 subcategories, and 3 core tasks . it covers 2D, 3D and animated plots across 5 widely used visualization libraries.
Fico: Evaluating Vision-Language Models under Visual Fidelity and Compression at Scale (2026.findings-acl)

Copied to clipboard

Challenge: Visual text compression is emerging paradigm for rendering text as images for processing by vision-language models.
Approach: They propose a benchmark to assess VLM robustness under dense visual inputs.
Outcome: Evaluating 13 general-purpose VLMs and 3 OCR-specialized models reveals performance drops sharply under increased density or reduced resolution; cross-task transfer between OCR, NIAH, and VQA is limited; and VQ is comparatively robust because low-level details are lost before high-level semantics.
Inject to Heal: Alleviating hallucination in LVLMs via Context Embedding Injection (2026.findings-acl)

Copied to clipboard

Challenge: a large vision-language model can generate hallucinations inconsistent with visual input . a lightweight method that embeds the last input token as a grounding signal reduces the likelihood of hallucinosity.
Approach: They propose a training-free mitigation strategy that harnesses the hidden state of the last input token as a grounding signal to maintain visual fidelity throughout decoding and curb hallucinations.
Outcome: The proposed method outperforms state-of-the-art methods on CHAIR, AMBER, and MMHal benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations